Papers with semantic similarity

49 papers
Multiplex Graph Neural Network for Extractive Text Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for extractive text summarization do not consider multiple types of inter-sentential relationships, nor model intra-sententential relationships.
Approach: They propose a novel method to combine different types of relationships among sentences and words to model sentence embedding.
Outcome: The proposed model is compared with existing methods on CNN/DailyMail benchmark dataset to demonstrate its effectiveness.
ALOHa: A New Measure for Hallucination in Captioning Models (2024.naacl-short)

Copied to clipboard

Challenge: Existing metric for object hallucination, CHAIR, is limited to MS COCO objects and synonyms.
Approach: They propose a new open-vocabulary metric, ALOHa, which leverages large language models to measure object hallucinations.
Outcome: The proposed metric correctly identifies 13.6% more hallucinated objects than CHAIR on HAT and 30.8% more on nocaps.
BeLLM: Backward Dependency Enhanced Large Language Model for Sentence Embeddings (2024.naacl-long)

Copied to clipboard

Challenge: Existing LLMs adopt autoregressive architectures without explicit backward dependency modeling.
Approach: They propose a backward dependency enhanced large language model that transforms attention layers from uni-to-bi-directional to learn sentence embeddings.
Outcome: The proposed model achieves state-of-the-art performance in varying scenarios.
What Makes Sentences Semantically Related? A Textual Relatedness Dataset and Empirical Study (2023.eacl-main)

Copied to clipboard

Challenge: Existing work on semantic relatedness has focused on semantic similarity because of a lack of relatedness datasets.
Approach: They propose a dataset for semantic relatedness that has 5,500 English sentence pairs manually annotated using a comparative annotation framework.
Outcome: The proposed dataset has 5,500 English sentence pairs manually annotated using a comparative annotation framework.
Leveraging Multi-lingual Positive Instances in Contrastive Learning to Improve Sentence Embedding (2024.eacl-long)

Copied to clipboard

Challenge: Recent trends in learning monolingual and multilingual sentence embeddings are based on contrastive learning (CL) among an anchor, one positive and multiple negative instances.
Approach: They propose to leverage multiple positives to improve learning of multilingual sentence embeddings by using an anchor, one positive, and multiple negative instances.
Outcome: The proposed approach improves retrieval, semantic similarity, and classification performance on unseen languages.
LAMP-MedQA: A Lightweight Multi-Agent System for Patient-Oriented Medical Question Answering (2026.acl-srw)

Copied to clipboard

Challenge: Large language models (LLMs) are a promising way to bridge the gap between patient health literacy and access to care.
Approach: They evaluate a range of open- and closed-source LLMs on a MeDiSumQA dataset . they propose a lightweight multi-agent framework for patient-oriented medical question answering .
Outcome: The proposed model achieves lower FKGL than zero-shot GPT-5 and highest simplification quality among all models.
Language-agnostic BERT Sentence Embedding (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for learning bilingual sentence embeddings are not well explored.
Approach: They propose to combine best methods for learning multilingual sentence embeddings with pre-trained models to achieve 83.7% bi-text retrieval accuracy over 112 languages on Tatoeba.
Outcome: The proposed model achieves 83.7% bi-text retrieval accuracy over 112 languages on Tatoeba, above the 65.5% achieved by LASER.
REMATCH: Robust and Efficient Matching of Local Knowledge Graphs to Improve Structural and Semantic Similarity (2024.findings-naacl)

Copied to clipboard

Challenge: Existing AMR metrics are inefficient and struggle to capture semantic similarity . Existing metrics are not efficient and lack a systematic evaluation benchmark .
Approach: They propose a new AMR similarity metric, rematch, which matches graphs structurally and semantically to each other.
Outcome: The proposed metric is five times faster than the next most efficient metric.
Interpretable Company Similarity with Sparse Autoencoders (2025.acl-industry)

Copied to clipboard

Challenge: Traditionally, company comparisons rely on relative returns and discrete classifications, or a combination of both.
Approach: They propose to use clusters of embeddings to enhance the interpretability of Large Language Models by decomposing Large Language models activations into interpretable features.
Outcome: The proposed clusters of embeddings capture the internal representation of a company description, rather than just semantic similarity alone.
Cross-lingual Text Classification with Heterogeneous Graph Neural Network (2021.acl-short)

Copied to clipboard

Challenge: Existing methods for cross-lingual text classification only consider factors beyond semantic similarity, causing performance degradation between some language pairs.
Approach: They propose a method to incorporate heterogeneous information within and across languages for cross-lingual text classification using graph convolutional networks.
Outcome: The proposed method significantly outperforms state-of-the-art models on all tasks and achieves consistent performance gain over baselines in low-resource settings.
TISE: A Tripartite In-context Selection Method for Event Argument Extraction (2024.naacl-long)

Copied to clipboard

Challenge: Recent studies show that LLMs can finish inference by providing several examples.
Approach: They propose a method which integrates three requirements when selecting an in-context example and integrates them into a set of determinantal point processes to enhance the reasoning capabilities of LLMs.
Outcome: The proposed method can achieve superior performance with fewer examples and outperform some supervised methods.
A Semantically Consistent and Syntactically Variational Encoder-Decoder Framework for Paraphrase Generation (2020.coling-main)

Copied to clipboard

Challenge: Paraphrase generation is a longstanding problem in natural language processing (NLP) Neural network-based methods have shown great progress on paraphrase generation.
Approach: They propose a framework that integrates variational inference on a target-related latent variable to introduce the diversity.
Outcome: The proposed framework outperforms baseline models on the metrics based on n-gram matching and semantic similarity, and it can generate multiple different paraphrases by assembling different syntactic variables.
ValCAT: Variable-Length Contextualized Adversarial Transformations Using Encoder-Decoder Language Model (2022.naacl-main)

Copied to clipboard

Challenge: Existing word-level approaches to attack text are limited to a single word . existing methods ignore interactions between consecutive words, resulting in one-to-one attacks .
Approach: They propose a black-box attack framework that misleads the language model by applying variable-length contextualized transformations to the original text.
Outcome: The proposed framework outperforms existing methods on classification and inference tasks.
GenSense: A Generalized Sense Retrofitting Model (C18-1)

Copied to clipboard

Challenge: Existing word embedding models use only one vector to represent a word, which is problematic in some natural language processing tasks that require sense level representation.
Approach: They propose a generalized sense embedding learning framework that integrates with the semantic relations between the senses, the relation strength and the semantic strength.
Outcome: The proposed model outperforms previous models in three types of experiments: semantic relatedness, contextual word similarity and semantic difference.
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)

Copied to clipboard

Challenge: SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text .
Approach: They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu.
Outcome: The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages.
Semantic Alignment with Calibrated Similarity for Multilingual Sentence Embedding (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for learning semantic similarity between two English sentences have focused on one sub-task and therefore showed biased performance.
Approach: They propose a method to learn semantic similarity between two English sentences using siamese networks.
Outcome: The proposed method improves on both sub-tasks and predicts similarity scores in 14 languages.
AdapterSoup: Weight Averaging to Improve Generalization of Pretrained Language Models (2023.findings-eacl)

Copied to clipboard

Challenge: Pretrained language models often need to specialize to specific domains.
Approach: They propose an approach that performs weight-space averaging of adapters trained on different domains.
Outcome: The proposed approach improves performance to new domains without extra training.
CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Existing natural language understanding benchmarks inadequately address the ability to evaluate causal relationships.
Approach: They propose to use CLEAR-3K to evaluate whether language models can determine if one statement causally explains another.
Outcome: The proposed questions show that language models often confuse semantic similarity with causality, relying on lexical and semantic overlap instead of inferring actual causal explanatory relationships.
CALAMR: Component ALignment for Abstract Meaning Representation (2024.lrec-main)

Copied to clipboard

Challenge: Abstract meaning representation (AMR) graphs represent semantic structure in a syntactic independent way.
Approach: They propose a method for graph alignment that can support summarization and evaluation.
Outcome: The proposed method produces graphs that explain what is summarized through their alignments, which can be used to train graph based summarization learners.
Aiming beyond the Obvious: Identifying Non-Obvious Cases in Semantic Similarity Datasets (P19-1)

Copied to clipboard

Challenge: Existing datasets for scoring text pairs in terms of semantic similarity contain instances whose resolution differs according to the degree of difficulty.
Approach: They propose to use lexical overlap to distinguish obvious from non-obvious text pairs by focusing on item difficulty and ground-truth labels to characterise existing datasets.
Outcome: The proposed models are based on lexical overlap and ground-truth labels and focus on cases of similarity which require more complex inference.
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that RLHF improves performance of Large Language Models . BT-based RMs struggle to distinguish between similar preference responses .
Approach: They propose to enhance BT-based reward models by using an adaptive margin mechanism . they use semantic similarity and reward-predicted reward differences to adjust focus .
Outcome: Experimental results show that the proposed method outperforms existing methods in both in-distribution and OOD settings.
RADAR: A Reasoning-Guided Attribution Framework for Explainable Visual Data Analysis (2026.findings-eacl)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) provide no visibility into which parts of visual data informed their conclusions.
Approach: They propose a semi-automatic approach to attribute reasoning process by highlighting regions in charts and graphs that justify model answers.
Outcome: The proposed method improves attribution accuracy by up to 15 percentage points compared to baseline methods and achieves high semantic similarity with ground truth responses.
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)

Copied to clipboard

Challenge: Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies.
Approach: They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances .
Outcome: The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations.
Making Fast Graph-based Algorithms with Graph Metric Embeddings (P19-1)

Copied to clipboard

Challenge: Graph measures, such as node distances, are inefficient to compute.
Approach: They propose a way to learn graph embeddings by using vector operations instead of a graph structure.
Outcome: The proposed method outperforms other graph embeddings on word similarity and word sense disambiguation tasks.
Learning Semantic Textual Similarity via Topic-informed Discrete Latent Variables (2022.emnlp-main)

Copied to clipboard

Challenge: Recent discrete latent variable models have received a surge of interest in both NLP and CV . they are comparable to the continuous counterparts in representation learning, but are more interpretable in their predictions.
Approach: They develop a topic-informed discrete latent variable model for semantic textual similarity . they inject the quantized representation into a transformer-based language model .
Outcome: The proposed model outperforms strong baselines in semantic textual similarity tasks.
Perceptual Models of Machine-Edited Text (2021.findings-acl)

Copied to clipboard

Challenge: a dataset of human judgments of machine-edited text is presented . we compare six different methods to create generic models of human perception .
Approach: They propose to use six machine-editing methods to model human perceptions of edited text . they use a dataset of human judgments of machine-edited text and scientific abstracts .
Outcome: The proposed model is based on human judgments of machine-edited text and scientific abstracts . human judgment of edited text is predicted to be within 6% of human consensus labeling .
Semantic Relatedness Based Re-ranker for Text Spotting (D19-1)

Copied to clipboard

Challenge: Existing approaches to text spotting are limited by semantic similarity, but they can be useful for other tasks.
Approach: They propose a neural approach to learn semantic relatedness from existing sentences.
Outcome: The proposed approach outperforms existing approaches when applied to a text spotting task.
RoMe: A Robust Metric for Evaluating Natural Language Generation (2022.acl-long)

Copied to clipboard

Challenge: Empirical results suggest that RoMe has a stronger correlation to human judgment over state-of-the-art metrics in evaluating system-generated sentences across several NLG tasks.
Approach: They propose an automatic evaluation metric incorporating several core aspects of natural language understanding (language competence, syntactic and semantic variation).
Outcome: The proposed evaluation metric is trained on language features such as semantic similarity combined with tree edit distance and grammatical acceptability, using a self-supervised neural network.
Beyond BLEU:Training Neural Machine Translation with Semantic Similarity (P19-1)

Copied to clipboard

Challenge: Recent work has shown that optimizing neural machine translation systems to directly improve evaluation metrics such as BLEU can improve final translation accuracy.
Approach: They propose a reward function that assigns partial credit to BLEU and provides more diversity in scores than BLUE.
Outcome: The proposed reward function improves translation accuracy, semantic similarity, and human evaluation on four languages trans-lated to English and the optimization procedure converges faster.
OpticE: A Coherence Theory-Based Model for Link Prediction (2022.coling-1)

Copied to clipboard

Challenge: Knowledge representation learning is a key step required for link prediction tasks with knowledge graphs (KGs).
Approach: They propose a new embedding approach based on the physical phenomenon of optical interference to reduce the semantic ambiguity in KGs.
Outcome: The proposed model can compete with existing methods on KG benchmarks.
Hard Emotion Test Evaluation Sets for Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Existing tests on emotion datasets do not show whether language models understand emotions or exploit supperficial lexical cues.
Approach: They propose to use two existing emotion datasets to evaluate whether language models make inferential decisions for emotion detection.
Outcome: The proposed test sets evaluate language models on emotion datasets.
ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-Translation (2023.acl-long)

Copied to clipboard

Challenge: Paraphrase generation is a long-standing task in natural language processing (NLP).
Approach: They propose to generate large-scale syntactically diverse paraphrase datasets by abstract meaning representation back-translation.
Outcome: The proposed dataset is syntactically more diverse than existing datasets while maintaining good semantic similarity.
Discovering Lobby-Parliamentarian Alignments through NLP (2024.naacl-long)

Copied to clipboard

Challenge: Influence of interest groups on parliamentarians and subversion of electorate to determine policy has led to demands from groups such as Transparency International .
Approach: They collect datasets of lobbies’ position papers and MEPs’ speeches and compare them on the basis of semantic similarity and entailment.
Outcome: The proposed method performs significantly better than baselines and matches the public meetings of MEPs with retweet links.
CLISTER : A Corpus for Semantic Textual Similarity in French Clinical Narratives (2022.lrec-1)

Copied to clipboard

Challenge: Modern Natural Language Processing relies on the availability of annotated corpora for training and evaluation.
Approach: They propose to annotate sentences in French using a definition of similarity guided by clinical facts and use it to evaluate the corpus.
Outcome: The proposed model can capture similarity with state-of-the-art performance on the DEFT STS shared task evaluation data set.
VAEGPT-Sim: Improving Sentence Representation with Limited Corpus Using Gradually-Denoising VAE (2024.findings-acl)

Copied to clipboard

Challenge: Text embedding requires a highly efficient method for training domain-specific models on limited corpora.
Approach: They propose a model that combines a denoising variational autoencoder with a target-specific discriminator to generate synonymous sentences that closely resemble human language.
Outcome: The proposed model surpasses ConSERT by 2.8 points in small-dataset training on STS benchmarks.
Urban Dictionary Embeddings for Slang NLP Applications (2020.lrec-1)

Copied to clipboard

Challenge: a new set of word embeddings is released to improve word embedment performance . word embeds provide useful representations of meanings of words in vectors .
Approach: They present a set of word embeddings trained on Urban Dictionary . they show they have high performance across a range of common word embeding evaluations .
Outcome: The first set of word embeddings trained on Urban Dictionary has high performance . the embeddables perform better on a range of common word evaluation tasks .
Watch Out Your Industrial Copilots: Stealthy Backdoor Attack Against LLM-Based PLC Code Generation (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are being used to generate PLC code from natural language.
Approach: They propose a stealthy backdoor attack framework targeting LLM-based PLC code generation . they incorporate six malicious logic injection patterns and a pipeline to refine stealthiness .
Outcome: The proposed framework achieves 82.92% success rate while remaining stealthy . it bypasses quality validation and is difficult to detect .
SemR-11: A Multi-Lingual Gold-Standard for Semantic Similarity and Relatedness for Eleven Languages (L18-1)

Copied to clipboard

Challenge: SemR-11 is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
Approach: This paper describes a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
Outcome: The dataset is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
AlignScore: Evaluating Factual Consistency with A Unified Alignment Function (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to evaluate factual consistency of text depend on limited data . e.g., generated text can contain factual inconsistencies that are irrelevant to context .
Approach: They propose a new holistic metric that measures factual inconsistencies . they use 4.7M training examples from 7 well-established tasks .
Outcome: The proposed metric outperforms existing metrics on 22 datasets and matches or outperFORMs them.
Semantic Similarity Covariance Matrix Shrinkage (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to estimate covariance matrix relied on historical price data and ignored company fundamental data.
Approach: They propose to use semantic similarity to improve covariance estimations by using a shrinkage target.
Outcome: The proposed method is compared with the prior art estimate for covariance shrinkage using semantic similarity and price history.
Beyond Contrastive Learning: A Variational Generative Model for Multilingual Retrieval (2023.acl-long)

Copied to clipboard

Challenge: Contrastive learning is the dominant paradigm for learning text representations from parallel text, but finding negative examples can be expensive in terms of compute or manual effort.
Approach: They propose a generative model for learning multilingual text embeddings which encourages source separation in multilingual contexts by an approximation.
Outcome: The proposed model outperforms both a strong contrastive and generative baseline on a suite of tasks including semantic similarity, bitext mining, and cross-lingual question retrieval.
On the Sentence Embeddings from Pre-trained Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Pre-trained contextual representations like BERT have been widely used for NLP tasks.
Approach: They propose to transform anisotropic sentence embedding distribution to smooth and isotropic Gaussian distribution by normalizing flows that are learned with an unsupervised objective.
Outcome: The proposed method achieves significant performance gains over state-of-the-art embeddings on a variety of semantic textual similarity tasks.
PaRaDe: Passage Ranking using Demonstrations with LLMs (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that large language models can be instructed to perform zero-shot passage re-ranking . Existing work like UPR demonstrate promising results for zero- shot ranking using LLMs .
Approach: They propose a demonstration selection strategy based on difficulty rather than semantic similarity . they propose to include only one demonstration in the prompt to improve re-ranking .
Outcome: The proposed method improves LLM-based re-ranking by adding one demonstration to the prompt.
Anonpsy: A Graph-Based Framework for Structure-Preserving De-identification of Psychiatric Narratives (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to de-identify narratives operate at the text level and offer limited control over which semantic elements are preserved or altered.
Approach: They propose a de-identification framework that reformulates the task as graph-guided semantic rewriting.
Outcome: The proposed framework preserves diagnostic fidelity while achieving low re-identification risk.
CROWD: Certified Robustness via Weight Distribution for Smoothed Classifiers against Backdoor Attack (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on the certification of robustness of NLP models against backdoor attacks have focused on empirical defences against adversarial attacks without formal guarantees.
Approach: They propose a model-agnostic mechanism for large-scale models that applies to complex model structures without the need for assessing model architecture or internal knowledge.
Outcome: The proposed model-agnostic mechanism is tested on a diverse range of datasets and tasks, showing it can be used to mitigat backdoor triggers.
CausalRAG: Integrating Causal Graphs into Retrieval-Augmented Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing RAG frameworks face critical limitations due to text chunking and semantic similarity.
Approach: They propose a framework that incorporates causal graphs into the retrieval process.
Outcome: The proposed framework preserves contextual continuity and improves retrieval precision, leading to more accurate and interpretable responses.
VecCISC: Improving Confidence-Informed Self-Consistency with Reasoning Trace Clustering and Candidate Answer Selection (2026.findings-acl)

Copied to clipboard

Challenge: Weighted majority voting requires a critic to evaluate each candidate’s reasoning trace to produce the answer’s confidence score.
Approach: They propose a lightweight framework that uses a measure of semantic similarity to filter reasoning traces that are semantically equivalent to others, degenerate, or hallucinated.
Outcome: The proposed framework reduces token usage by 47% while maintaining or exceeding the accuracy of CISC.
DeepSpecs: Expert-Level Question Answering in 5G (2026.findings-acl)

Copied to clipboard

Challenge: 3GPP standards define the technical design and implementation of 5G systems . expert-level questions require navigating thousands of pages of cross-referenced standards .
Approach: They propose a standard-native retrieval-augmented generation system that can answer 5G questions . they use SpecDB, ChangeDB, TDocDB and a metadata-rich retrieval system to do this .
Outcome: The proposed solution outperforms base models and state-of-the-art RAG systems in QA datasets . expert-level queries require navigating thousands of pages of cross-referenced standards .
From Regulatory Approvals to Patents: Cross-Domain Linking for Cardiovascular Device Traceability (2026.acl-long)

Copied to clipboard

Challenge: a new cross-domain entity linking problem exists between FDA-approved medical devices and patents . a recent study compared the semantics of FDA documents with patents, resulting in minimal overlap .
Approach: They propose a framework that links FDA-approved medical devices to their patents . they propose 'MedDevKG' framework that integrates a domain-specific ontology .
Outcome: The proposed framework outperforms existing methods in lower-bound recalls and noise reductions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations